Papers with data creation
FAMULUS: Interactive Annotation and Feedback Generation for Teaching Diagnostic Reasoning (D19-3)
Copied to clipboard
Jonas Pfeiffer, Christian M. Meyer, Claudia Schulz, Jan Kiesewetter, Jan Zottmann, Michael Sailer, Elisabeth Bauer, Frank Fischer, Martin R. Fischer, Iryna Gurevych
| Challenge: | Existing systems for technologyenhanced learning address skills on recalling, explaining, and applying knowledge, e.g., in automatically generated language learning exercises and math word problems. |
| Approach: | They propose to leverage a NLP model to support experts in their further data annotation with automatic suggestions and provide automatic feedback for students. |
| Outcome: | The proposed system improves on two user studies on diagnostic reasoning in medicine and teacher education and can be extended to further use cases. |
All You May Need for VQA are Image Captions (2022.naacl-main)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. |
| Approach: | They propose a method that automatically derives VQA examples at volume by leveraging existing image-caption annotations combined with neural models for textual question generation. |
| Outcome: | The proposed method improves state-of-the-art zero-shot accuracy by double digits and achieves robustness that lacks in the same model trained on human-annotated VQA data. |
WikiNEuRal: Combined Neural and Knowledge-based Silver Data Creation for Multilingual NER (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a key intermediate task in NLP. |
| Approach: | They propose a method which uses knowledge-based approaches and neural models to produce high-quality training corpora for NER. |
| Outcome: | The proposed method improves on standard benchmarks and yields significant improvements up to 6 span-based F1-score points over previous state-of-the-art systems for data creation. |
UniSumEval: Towards Unified, Fine-grained, Multi-dimensional Summarization Evaluation for LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks for summarization quality evaluation lack diverse input scenarios, focus on narrowly defined dimensions, and struggle with subjective and coarse-grained annotation schemes. |
| Approach: | They propose to use AI to help human annotations and identifie potentially hallucinogenic input texts. |
| Outcome: | The proposed benchmarks improve on existing benchmarks in terms of input diversity, granularity of human annotations, and evaluation dimensions. |
Analysis of Automatic Annotation Suggestions for Hard Discourse-Level Tasks in Expert Domains (P19-1)
Copied to clipboard
Claudia Schulz, Christian M. Meyer, Jan Kiesewetter, Michael Sailer, Elisabeth Bauer, Martin R. Fischer, Frank Fischer, Iryna Gurevych
| Challenge: | Existing deep learning methods require large amounts of training data to achieve reasonable performance. |
| Approach: | They propose to generate automatic annotation suggestions for a discourse-level sequence labelling task that requires extensive domain expertise. |
| Outcome: | The proposed model improves with newly annotated texts while introducing no biases. |
IMPARA: Impact-Based Metric for GEC Using Parallel Data (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for automatic evaluation of grammatical error correction require multiple reference sentences or manual scores. |
| Approach: | They propose an Impact-based Metric for GEC using PARAllel data, IMPARA . IMPRA computes correction impacts computed by parallel data comprising pairs of grammatical/ungrammatically-spaced sentences. |
| Outcome: | The proposed method can perform evaluations that fit different domains and correction styles. |
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search (2025.acl-long)
Copied to clipboard
Linhao Yu, Xingguang Ji, Yahui Liu, Fanheng Kong, Chenxi Sun, Jingyuan Zhang, Hongzhi Zhang, V. W., Fuzheng Zhang, Deyi Xiong
| Challenge: | Existing benchmarks and evaluation protocols suffer from inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes. |
| Approach: | They propose an automatic framework which leverages Monte Carlo Tree Search to construct numerous and diverse descriptive sentences that thoroughly represent video content in an iterative way. |
| Outcome: | The proposed framework improves MCTS-VCB and DREAM-1K on video captioning tasks by 25.0% and 16.3% respectively. |
K-UniMorph: Korean Universal Morphology and its Feature Schema (2023.findings-acl)
Copied to clipboard
| Challenge: | Previously, the Korean language has been underrepresented in the field of morphological paradigms amongst hundreds of diverse world languages. |
| Approach: | They propose a new Universal Morphology dataset for Korean that preserves its distinct characteristics. |
| Outcome: | The proposed dataset extracts inflected Korean verb forms from the largest annotated corpus for Korean. |
MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution (2026.acl-long)
Copied to clipboard
Jianwen Chen, Xinyu Yang, Peng Xia, Arian Azarang, Yueh Z Lee, Gang Li, Hongtu Zhu, Yun Li, Beidi Chen, Huaxiu Yao
| Challenge: | Recent advances in large reasoning models have broadened the capabilities of medical artificial intelligence. |
| Approach: | They propose a reasoning framework for complex medical inference that reformulates medical reasoning as a parallelizable directed acyclic graph process based on Petri Net theory. |
| Outcome: | The proposed reasoning framework improves strong general-purpose LLMs by up to 8.9%. |
VIMQA: A Vietnamese Dataset for Advanced Reasoning and Explainable Multi-hop Question Answering (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing Vietnamese Question Answering (QA) datasets do not explore the model’s ability to perform advanced reasoning and provide evidence to explain the answer. |
| Approach: | They propose to use Vietnamese as a question-answer dataset with 10,000 Wikipedia-based multi-hop question-and-answ pairs to test model's ability to reason and explain the answer. |
| Outcome: | The proposed dataset is in Vietnamese, a low-resource language. |
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages. |
| Approach: | They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks. |
| Outcome: | The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language . |